Back

Nature Biotechnology

Springer Science and Business Media LLC

Preprints posted in the last 7 days, ranked by how well they match Nature Biotechnology's content profile, based on 172 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit.

1
Calibration-free compression brings Evo 2 to its full million-token context on a single GPU

Patsakis, M.; Tzanakakis, A.; Georgakopoulos-Soares, I.

2026-09-01 bioinformatics 10.64898/2026.08.28.747902 medRxiv
Top 0.1%
15.3%
Show abstract

Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2s weights and attention cache to four bits without calibration data, and serves both through fused kernels. Compression is near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant variant-effect prediction. It brings Evo 2 40B onto one 80 GB GPU and Evo 2 7B to its full million-token context within a 40 GB memory budget, an eightfold gain in reachable context. We further show that the released chunked-prefill path is silently incorrect, returning plausible but uncorrelated likelihoods, and derive the block-wise continuation that repairs it: a complete 580-kilobase bacterial genome is now scored in one context in 22 minutes rather than 13.7 hours.

2
Data-driven spectroscopic dictionaries and detector-calibrated inference for photon-limited Raman hyperspectral imaging of living cells

Yagi, S.; Sagami, N.; Eshima, I.; Hiramatsu, K.

2026-09-01 cell biology 10.64898/2026.08.31.748229 medRxiv
Top 0.1%
11.8%
Show abstract

Label-free Raman imaging of living cells is photon limited: at exposures compatible with cellular dynamics, single-pixel spectra carry about one count per channel on a dominant smooth background. We present an unmixing framework in which the decoder of a physics-constrained autoencoder is restricted to a data-driven spectroscopic dictionary: band centers,widths, and pseudo-Voigt shapes are measured from the dataset and fixed, and the network learns only nonnegative band amplitudes, a smooth B-spline background, and a per-pixel gain.First, on slit-scanning images of HeLa cells (532 nm) the dictionary yields spike-free component spectra that read as band tables, including a resonance-enhanced cytochrome-c-associated component matching literature spectra, and the most stable decomposition against the component number. Second, the dictionary and initialization calibrated at 1 s exposure perline transfer to 100 ms per line (12 s sweeps): cytochrome-c spectral identity survives a single sweep (correlation 0.92) while its map remains photon limited; the dictionary provides spectral physicality, and the transferred initialization prevents a structural collapse that global map correlations miss; in a measurement-derived phantom the dictionary estimator holds thecytochrome-c spectrum to 17-19{degrees} spectral angle at 100 ms, where classical factorizations and free decoders lose it (55-64{degrees}). Estimation on the count-equivalent detector output uses a calibrated shifted-Poisson quasi-likelihood. Third, evaluation must be time matched:correlation against a separately acquired reference saturates through slow specimen drift and acquisition mismatch rather than photon noise, and the self-consistency of learned denoisers is inflated by shared bias; time-matched self-consistency and independent cross-checks areproposed.

3
Vipsania: Unsupervised Deep Gene Finding

Krieg, R.; Becker, F.; Saenko, S.; Diehl, J.; Stanke, M.

2026-08-30 bioinformatics 10.64898/2026.08.26.747235 medRxiv
Top 0.1%
11.6%
Show abstract

Scaling the structural annotation of protein-coding genes to all eukaryotic genomes remains a major challenge. While recent deep learning methods rival evidence-based pipelines without requiring RNA-seq or alignments, they are entirely supervised. They depend on large, high-quality training sets from diverse genomes, leaving many basal eukaryotic clades without an accurate ab initio gene finder. We present Vipsania, the first unsupervised deep gene finder. A differentiable hidden Markov layer inside a deep sequence model learns to predict gene structures from unannotated genomes alone. Vipsania is pretrained for virtually all eukaryotes and finetunes without supervision on the target genome. It is, on average, more accurate than supervised methods across most clades and avoids the accuracy drop that supervised models suffer on distant target genomes. Vipsania adapts to non-standard genetic codes and provides a fast and highly versatile tool for unbiased, pan-eukaryotic genome annotation. The source code is available at https://github.com/gaius-augustus/vipsania.

4
Aerolysin enables modular, non-genetic functionalization of living cell surfaces

Lemmex, A. C.; Pawlak, M. R.; Gordon, W. R.

2026-08-31 biochemistry 10.64898/2026.08.28.746739 medRxiv
Top 0.4%
6.6%
Show abstract

Methods for installing synthetic functions on living cell surfaces provide powerful approaches for imaging, sensing, and manipulating cell behavior, but many require genetic modification of the target cell or chemical modification of the plasma membrane. Here, we repurpose the glycosylphosphatidylinositol-anchored protein (GPI-AP)-binding toxin aerolysin as a modular chassis for non-genetic cell-surface functionalization. We show that a non-cytotoxic, monomeric aerolysin mutant retains high-affinity and GPI-AP-dependent cell binding when genetically fused to diverse protein cargos. Fluorescent protein-aerolysin fusions robustly label multiple cell types and remain predominantly associated with the cell surface for at least 24 h, in contrast to wheat germ agglutinin, which is extensively internalized. Aerolysin can also be equipped with SpyTag/SpyCatcher to enable modular assembly with independently expressed protein cargos. Importantly, aerolysin supports functional rather than solely optical modification of the cell surface: fusion to the proximity-labeling enzyme APEX2 enables extracellular protein biotinylation, while fusion to HUH endonuclease tags enables covalent attachment of synthetic DNA to living cells. Using this latter architecture, we developed a DNA hairpin sensor that converts cell-surface nuclease activity into a fluorescent signal and distinguishes cells with different levels of extracellular nuclease activity. Together, these results establish non-cytotoxic aerolysin as a genetically encoded, soluble adapter for installing proteins, enzymes, and programmable nucleic acids onto living cells without modification of the target-cell genome.

5
RegimeFormer: A Large Protein Model of Global Perturbation Regimes

Ma, S.; Chai, Y.; Wu, Y.; Zhang, Q.; Yuan, Y.; Zhao, K.; Chen, Z.; Wang, H.; Cao, S.; Yu, X.; Han, X.; Liu, Y.; Liu, Y.; Zhu, T.; Tao, D.

2026-08-30 bioinformatics 10.64898/2026.08.26.747182 medRxiv
Top 0.5%
6.6%
Show abstract

Protein language models organize sequence and structure at scale, but a global representation of how proteins respond to mutation remains lacking. We present RegimeFormer, a large protein perturbation model coupled to RegimeAtlas, constructed by harmonizing and indexing 202,556,313 non-redundant protein sequences across the tree of life. A diversity-preserving one-million-protein subset provides the high-resolution training and inference layer, with 995,995 proteins yielding residue-level summaries across 407,048,356 residues and substitution-specific predictions available on demand. Across experimental deep mutational scanning, molecular benchmarks, structural confidence and evolutionary constraint, RegimeFormer identifies reproducible protein-level perturbation regimes that organize residue fragility, adaptability and predictive uncertainty. Regime conditioning improves substitution-specific prediction, with the largest relative gains under unseen-protein, unseen-family and low-homology evaluation. RegimeFormer-derived molecular priors further improve downstream transcriptomic and drug-response modelling. Together, RegimeFormer and RegimeAtlas provide a scalable framework for mapping, predicting and querying protein perturbation landscapes across global sequence space.

6
RECON infers regions of interest from H&E images and reconstructs whole-slide molecular profiles at single-cell resolution

Yang, X.; Hao, N.; Zhao, R.; Angel, S.; Tan, Y.; Lian, C. G.; Zhou, L.; Olson, D.; Yu, K.-H.; Ruiz de Luzuriaga, A.; Wan, G.

2026-09-01 bioinformatics 10.64898/2026.08.25.747122 medRxiv
Top 0.5%
6.2%
Show abstract

Spatial omics technologies resolve molecular expression and spatial architecture at single-cell resolution, but profiling whole slides remains costly. In practice, only a few regions of interest (ROIs) are profiled, leaving the rest of the tissue unmeasured. S2-omics was the first framework to unify ROI selection with out-of-ROI prediction, but it operates on superpixels rather than individual cells and predicts discrete cell types rather than continuous molecular profiles. Superpixel-based representations do not explicitly preserve cell boundaries, while categorical cell-type labels cannot quantify molecular expression within cells. Here we present RECON, a two-stage framework that performs ROI inference and whole-slide molecular reconstruction at single-cell resolution, predicting both continuous molecular profiles and discrete cell-type labels. In the first stage, RECON extracts morphological and microenvironmental features from individual cells to identify a representative ROI for spatially resolved single-cell molecular profiling. In the second stage, RECON trains deep learning models on molecular measurements acquired within the selected ROI and reconstructs transcriptomic or proteomic profiles for all remaining cells on the slide. Benchmarked against pathologist annotations, RECONs ROI selection outperforms the superpixel-based S2-omics approaches (IoU: 0.75 versus 0.64). For transcriptomics, refining the modeling unit from superpixels to single cells improves per-gene Pearson correlation by 22%. For proteomics, RECON surpasses the current state-of-the-art method, ROSIE, across all 16 markers, with a median per-cell Pearson correlation of 0.91 versus 0.84. Moreover, RECON delineates tumour boundaries and regions with distinct immune-cell densities, and highlights candidate tertiary lymphoid structures. Together, these results demonstrate that RECON enables informative ROI selection and whole-slide molecular reconstruction at single-cell resolution for both spatial transcriptomics and spatial proteomics.

7
ChemIntelligence Enables Antibody-Free, Ultra-Low-Input Profiling of Lysine Lactylation and Diverse Acyl-Proteomes

Shao, C.; He, Z.; Yuan, Q.; Giurcoiu, V.-G.; He, X.; Cao, X.; Huang, H.; Zhang, Y.; Zhang, Y.; Wang, D.; Jiang, Q.; Guo, Z.; Hao, H.; Wilhelm, M.; Ye, H.

2026-08-31 biochemistry 10.64898/2026.08.28.746934 medRxiv
Top 0.9%
4.3%
Show abstract

Lysine acylations, including lactylation (Klac), are pivotal regulators of cellular physiology. However, their analysis is currently bottlenecked by antibody enrichment strategies that suffer from sequence bias and require milligram-scale protein inputs, severely precluding the profiling of scarce clinical biopsies and rare cell populations. Here we present ChemIntelligence, an acyl-NHS chemistry-empowered derivatization strategy that rapidly generates unprecedented acylation-specific spectral libraries, exemplified by over 2.5x10^9 human Klac peptides, enabling cross-species reference atlases. Integrated with Prosit-based rescoring, these libraries substantially increase Klac identifications across diverse proteomic datasets. Leveraging this spectral resource, we devised ChemIntelligence Scope, a reproducible, multiplexed parallel reaction monitoring (PRM) platform that quantifies hundreds of Klac peptides per injection from as little as ~200 ng of cell lysates, clinical biopsies, and even true single cells - revealing functional Klac signatures inaccessible to conventional methods. The ChemIntelligence pipeline also extends seamlessly to lysine nicotinylation, underscoring its broad adaptability for discovering and profiling new acylations. Together, these chemical and computational advances establish a scalable, antibody-free framework for acyl-proteome mapping that overcomes input constraints and enables deep functional insights from otherwise intractable biological samples.

8
Genome-scale label-free imaging reveals cellular physiology encoded in bacterial collective architecture

Mellick, S. N. S.; Derringer, J. J.; Boyes, D.; Croteau, G.; Burke, M.; Gifford, S.; Stark, D. J.; Mike, L. A.; Turecki, S.; Carja, O.; Mikheyeva-Bridges, I. V.; Bridges, D. A.

2026-08-31 microbiology 10.64898/2026.08.30.748126 medRxiv
Top 1%
4.1%
Show abstract

DNA sequencing unified microbial genotyping into a single, comprehensive readout, yet phenotyping remains a slow and fragmented endeavor. Here, we introduce Microbial Phenotyping Using Low-magnification Label-free Imaging (PULLI), a computer vision platform that extracts microcolony and population-level phenotypes from brightfield timelapses of liquid culture growth. Using PULLI, we screened a genome-scale Vibrio cholerae mutant library, recording more than 200,000 images, which revealed that core bacterial pathways shape community architecture. Functionally related mutants converge in appearance, allowing us to resolve processes as distinct as biofilm formation, motility, central metabolism, cofactor biosynthesis, and envelope composition using a single approach. We further show PULLI can be used to determine a drug target, characterize other pathogens, and classify bacterial species. Our results show that bacterial multicellular development is an interpretable signature of genotype-phenotype relationships, which can be captured from simple brightfield timelapses. We release the PULLI pipeline and an interactive atlas of community forms.

9
PhageTAILor leverages machine learning for phage tail-like elements detection and classification in plant-associated bacteria

Cho, H.; Hour, S.; Roux, S.; Coclet, C.; Amusat, O.; Mutalik, V. K.; Kazakov, A. E.; Levy, A.; Nachmias, N.; Aureli, L.; Sweet, T. S.; Visel, A.; Ceballos, R. M.; Basso, J. T. R.

2026-09-01 microbiology 10.64898/2026.08.24.746745 medRxiv
Top 1%
3.6%
Show abstract

Phage tail-like elements (PTEs) -- tailocins, bacterial type VI secretion systems (T6SS), and extracellular contractile injection systems (eCIS) -- are contractile nanomachines that bacteria use to kill their neighbors and compete within their micro-ecosystems. PTEs help shape microbial community composition. Most PTE detection tools only detect a single PTE class. Moreover, most tailocin detection methods are largely restricted to Pseudomonas, leaving a key part of tailocin diversity uncharacterized. In this work, we present PhageTAILor (https://github.com/hjcho-bio/PhageTAILor), an integrative and fully automated pipeline that detects and classifies prophages and 3 PTE classes from bacterial genomes. PhageTAILor combines a 6-detector homology-based candidate search (geNomad, tail-gene, PHROGs-tail, SecReT6, eCIStem, and a divergence-tolerant tail-HMM detector) with a LightGBM classifier comprising 1 multiclass and 3 binary heads, trained on 6,501 bacterial genomes carrying 13,082 prophages and PTEs. A phylogeny-free feature matrix used in our model keeps predictions reproducible between model construction and user inference. PhageTAILor performs strongly at the genome level and generalizes beyond its Pseudomonas-rich training set. On a 76-strain cross-clade benchmark, PhageTAILor detected tailocins at F1 = 0.955. Furthermore, it identified 12 of 13 experimentally validated tailocins spanning five genera versus 2 of 13 for a Pseudomonas-restricted tool TattleTail. PhageTAILor also demonstrated sensitivity equivalent to viral detection tool geNomad while avoiding its higher false-positive rate. Applied to 7,925 plant- and soil-associated bacterial isolates, PhageTAILor showed that prophages in the phyllosphere and tailocins in plant-associated bacteria, whereas eCIS are enriched in soil. PhageTAILor is distributed as an open-source, modular pipeline with a command-line interface.

10
MechanoMaST - a multimodal pipeline for spatially registering mechanical and transcriptomic tissue data

Decker, L.; Olisov, D.; Schleussner, N.; Wiethoff, H.; Schmidt, T.; Nienhueser, H.; Pausch, T. M.; Korbel, J. O.; Diz-Munoz, A.

2026-08-31 biophysics 10.64898/2026.08.29.747727 medRxiv
Top 1%
3.5%
Show abstract

Spatial-omics workflows enable molecular analysis within tissue spatial context. Despite the prognostic value of tissue stiffness, these approaches have not incorporated direct, mechanical measurements. This omission reflects several challenges, including sample requirements, low throughput, specialized equipment, and complex data registration. Here, we introduce mechanoMaST (mechanics mapped to spatial transcriptomics), the first workflow to combine absolute mechanical measurements with spatial-omics. It pairs atomic force microscopy-based nanoindentation stiffness maps with spatial transcriptomics maps from adjacent tissue cryosections. The two modalities are then computationally co-registered to enable direct spatial correlation at 100 um resolution, with mapping accuracy quantified through error propagation, providing ground-truth mechanical data directly linked to spatial gene expression. We demonstrate mechanoMaST in human colorectal cancer liver metastasis, generating a spatial resource from 10 patients and revealing a four-gene stiffness signature. mechanoMaST is readily adaptable to other tissues across development and disease, and extendable to additional spatial-omics modalities in adjacent sections.

11
scPyviewer: a Python-native interactive viewer from AnnData single-cell data

Xuan, H.; Huang, Y.; Bian, J.; Liu, X.

2026-08-31 bioinformatics 10.64898/2026.08.26.747418 medRxiv
Top 1%
3.4%
Show abstract

Motivation: Interactive tools that let non-programmers explore an analyzed single-cell dataset, its embeddings, gene expression, cell metadata, and marker genes, have become standard laboratory infrastructure. Every actively maintained tool in this space (ShinyCell, ScRDAVis, sCIRCLE, scViewer) is built on R Shiny and requires a Seurat object as input. Laboratories whose primary analysis pipeline is Python/scanpy, the dominant framework for single-cell RNA-seq, spatial, and multi-omic analysis, therefore have no lightweight, language-native option that pairs a shareable web-based viewer with a scriptable Python API: sharing a scanpy result means either exporting to Seurat first or handing over a notebook that only a programmer can run. Results: We present scPyviewer, a web-based viewer that ingests AnnData objects directly and reproduces the core interaction patterns of the incumbent R Shiny tools without leaving the Python stack. In a feature-parity audit against three actively maintained R Shiny incumbents, scPyviewer matches or exceeds every baseline capability (7/7); among these, it uniquely offers native AnnData ingestion with no Seurat conversion, and cross-dataset comparison over shared genes and matched cell-type composition. Benchmarked head-to-head against the R/Seurat rendering substrate the incumbents are built on, identical operations, identical data, across three datasets spanning 22,315 to roughly 313,000 cells, scPyviewer renders every core view faster at every scale tested (up to 3.6x on a single view) and at a fraction of the memory (5.2x lower on the smallest dataset). At the largest scale tested, the gap becomes categorical rather than incremental: scPyviewer completes every view on a 313,000-cell dataset while the Seurat substrate exhausts an 8 GB memory budget and fails outright. Beyond the interactive app, scPyviewer installs via pip or conda and exposes a public Python API that returns Matplotlib figures and pandas tables for scripted, publication-ready output. Availability and implementation: scPyviewer is implemented in Python 3.11 (scanpy 1.11.5, anndata 0.12.19, streamlit 1.59.2, plotly 6.9.0) and distributed with a one-command reproduction interface that installs pinned dependencies, regenerates the benchmark and all figures, and launches the interactive app. Source code is available at https://github.com/xuan13hao/scPyviewer.git.

12
Single-Cell Inference of Structural States Of Ribosomes

Joly-Smith, E.; VanInsberghe, M.; Sarieva, K.; Marinelli, E.; van Es, R. M.; Sobrevals Alcaraz, P.; Vos, H. R.; Andersson-Rolf, A.; Clevers, H.; van Oudenaarden, A.

2026-08-31 molecular biology 10.64898/2026.08.29.747780 medRxiv
Top 1%
3.4%
Show abstract

Protein synthesis is dynamically regulated to control cell growth, differentiation, and stress responses. Recent single-cell sequencing methods can map ribosome positions on individual transcripts, but cannot capture the global translational states that coordinate protein synthesis across the transcriptome. In contrast, methods that measure the global translational landscape, such as polysome profiling and cryogenic electron tomography, lack either single-cell resolution or throughput. Here we introduce SCISSOR (Single-Cell Inference of Structural States of Ribosomes), a strategy that infers global translation activity in individual cells from the differential protection of ribosomal RNA (rRNA) against nuclease digestion. By integrating these protection signatures with the structure of the ribosome, SCISSOR resolves multiple ribosomal states and quantifies their abundance across thousands of individual cells. Applying SCISSOR reveals systematic variation in global translation across the cell cycle in human cells, as well as during the differentiation of murine intestinal stem cells into distinct epithelial lineages. These findings uncover principles of global translational regulation that are invisible to transcriptomic or ribosome-profiling assays, establishing a framework for studying global translation control at single-cell resolution.

13
Whole-body Super-resolution Functional and Molecular Imaging with Panoramic Photoacoustic-Ultrasound Tomography

Yao, R.; Husain, I.; Luo, J.; Huo, H.; Cai, X.; Wang, N.; Vu, T.; Li, J.; Xu, Y.; Menozzi, L.; Yang, J. J.; Lowerison, M.; Luo, X.; Song, P.; Yao, J.

2026-09-01 bioengineering 10.64898/2026.08.28.747673 medRxiv
Top 1%
3.3%
Show abstract

Photoacoustic (PA) and ultrasound (US) imaging provide complementary molecular, functional, and anatomical contrasts. Here, we present a panoramic PA-US imaging platform that integrates multispectral PA computed tomography (PACT) along with reflection-mode and transmission-mode US imaging through a single shared full-ring ultrasound array. We employ an ultrafast planewave transmission scheme in reflection-mode US for power Doppler (PWD) imaging and ultrasound localization microscopy (ULM). Additionally, we use the transmission-mode US to reconstruct a spatially resolved speed of sound (SoS) map that corrects both PA and US reconstruction. Such correction sharpens the resolution of PACT, suppresses the artifacts of PWD, and improves microbubble localization of ULM. Elevational scanning further enables whole-body volumetric imaging with co-registered PA and US contrasts. The integrated system maps photoswitchable DrBphP1-expressing tumors alongside their blood perfusion and oxygenation environment. Applying the platform to monitor unilateral renal ischemia-reperfusion injury, we report that microvascular perfusion and renal oxygenation recover at different rates. Collectively, we demonstrate that the integrated PA-US imaging platform provides a unified framework for multiparametric study of anatomy, perfusion, microvascular flow, oxygenation, and molecular activities.

14
Audited vibe coding suggests partial fetal-like convergence of tumor proteomes

Meyer, J. G.

2026-08-31 cancer biology 10.64898/2026.08.26.745609 medRxiv
Top 1%
3.2%
Show abstract

The balance between how much human tumors recapitulate fetal tissue programs versus lose adult tissue identity remains unresolved. I used audited vibe coding, a human-mediated, cross-model critique-and-refinement workflow, to re-analyze a public pan-cancer proteomic atlas. A primary large language model wrote and executed the analysis under scientific direction, while a separate model family audited the code, outputs and claims; findings were returned for correction across seven versioned releases. Among 229 tumor-adjacent pairs in seven organs, tumor-minus-adjacent proteomic change partially aligned with reverse fetal-to-adult maturation (organ-balanced cosine, 0.240; 95% interval, 0.138 to 0.335), with positive alignment in 189 of 229 patients (82.5%). The organ-balanced projection coefficient was 0.195 (95% interval, 0.069 to 0.244), indicating movement along only part of the developmental distance. Although reverse maturation overlapped adult-identity loss, a positive developmental component remained after identity loss entered first (0.203; 95% interval, 0.129 to 0.239). Suppression of adult-high proteins contributed to more positive alignment than reactivation of fetal-high proteins. The vibe coding audits identified substantive defects. A common-mask correction reduced the matched-organ advantage from 0.074 to 0.059; a missing-value correction barely changed aggregate geometry but replaced 5 of the top 40 liver contributors; and coupled resampling repaired uncertainty accounting without changing patient scores. As with any single report, the "vibe reanalysis" biological results are candidate discoveries pending independent replication. The workflow is a single feasibility case, not a reliability benchmark, and shows how conversationally generated analysis can be made more inspectable when model-written code is treated as untrusted, versioned and subject to separate-model critique and executable checks.

15
Chemi-Proteome Language Attention Network Empowers Fragment-Based Ligand Interactome and Binding Sites Discovery with Evidence

Liao, B.; He, J.; zhao, M.; Cui, X.; Cui, Y.; Dong, C.; Sun, H.; Zhang, L.; Zhang, J.

2026-08-30 bioinformatics 10.64898/2026.08.26.747036 medRxiv
Top 1%
3.2%
Show abstract

Deep learning has accelerated drug discovery, yet most existing models are trained using in vitro affinity datasets and consequently remain disconnected from the cellular context in which functional ligand-protein interactions occur. This limitation hinders the ability to reflect the complexity of native interactomes and characterize biological responses to molecular perturbation. Here we introduce C-PLANK (Chemi-Proteome Language Attention NetworK), a deep learning framework trained on fragment-protein interactions profiled directly in living cells using fully functionalized fragment (FFF) chemoproteomics. C-PLANK combines physicochemical embeddings with a bilinear attention network (BAN) to model both global cellular context and local residue-atom interactions, generating interpretable interaction fingerprints. Particularly, C-PLANK incorporates Cellular Interaction State Index (CISI), a systems-level evidential metric that contextualizes the biological plausibility of each predicted interaction against the global cellular interaction landscape. Across 431 ligand interactomes curated from eight independent chemoproteomic studies, C-PLANK consistently outperformed current state-of-the-art interaction prediction frameworks under both random and cold-protein evaluation settings. The inferred interaction fingerprints aligned with orthogonal evidence from structure-based pocket predictions, co-crystal structures, and cellular binding-site annotations. C-PLANK further generalized to unseen ligands. In a cellular target-focused discovery campaign, C-PLANK identified a previously unrecognized ligand that was subsequently advanced into an active chemical probe acting as a SIRT3 agonist in cellular assays. By learning directly from cellular chemoproteomics, C-PLANK moves beyond isolated interaction prediction toward cellular interaction-state modelling, establishing a computational foundation for future digital-twin frameworks in drug discovery.

16
Lineage-specific X chromosome inactivation escape and skew underlie sex-biased immune gene dosage and deleterious variant exposure

Kavanagh, D.; Steel, A.; King, H. E.; Vieira, H. G. S.; Kumar, K. R.; Masle-Farquhar, E.; King, C.; Skvortsova, K.; Weatheritt, R. J.

2026-08-31 bioinformatics 10.64898/2026.08.26.739472 medRxiv
Top 1%
3.1%
Show abstract

The X chromosome carries an unusually high density of immune genes and is a major contributor to sex differences in immune function and autoimmune diseases. In females, X-chromosome inactivation (XCI) has two major functional consequences: it shapes X-linked gene dosage through XCI escape and determines the cellular exposure of heterozygous X-linked variants through XCI skew. Yet because XCI creates a mosaic of cells expressing different parental X chromosomes, these properties have remained largely inaccessible in individual women, becoming measurable only where XCI is non-random or after aggregation across large cohorts. Consequently, how X-linked variation contributes to sex-biased immunity and differs between individual women has remained unresolved. Here we present scDaisyChain, a graph-based framework that reconstructs chromosome-scale X haplotypes directly from heterozygous SNPs and single-cell long-read transcriptomes. scDaisyChain achieves near-ground-truth accuracy in highly polymorphic mouse hybrids and shows strong concordance with orthogonal long-read whole-genome phasing in human samples. Applied to peripheral blood immune cells from healthy women, it reveals a lineage-specific escape program in which lymphoid cells escape XCI more broadly than monocytes, with corresponding gains in the inactive X chromatin accessibility and female-biased expression. Lineage-specific skew further alters the proportion of cells expressing each heterozygous X-linked variant, a property we term variant exposure. Predicted deleterious variants are preferentially found in low-exposure states, exemplified by a splice-altering TLR8 variant expressed in few cytotoxic T cells. In rheumatoid arthritis (RA), the monocyte compartment - which has the lowest escape in health - shows reproducible inactive X dysregulation converging on a trained-immunity programme linked to disease flare and synovial macrophage activation, with elevated escape of IL13RA1 and HDAC8. These findings establish lineage-specific escape, skew and variant exposure as quantifiable, patient-resolved determinants of sex-biased immune gene dosage and X-linked variant penetrance in health and autoimmune disease, resolving a dimension of female biology that has been previously inaccessible in individual donors.

17
In vivo multimodal lineage tracing of mammalian development by DeepTrack barcoding

Guo, C.; Jiang, J.; Wang, X.; Huang, X.; Zhang, S.; Shao, C.; Zhang, M.; Hu, X.; Yang, W.; Shang, F.; Wang, X.; Zhai, H.; Du, Q.; Liu, F.; He, D.; Liu, X.; Peng, G.; Cheng, S.; Zhang, Y.; Pei, D.; Pei, W.

2026-08-31 developmental biology 10.64898/2026.08.29.748052 medRxiv
Top 2%
2.7%
Show abstract

A comprehensive recording of cell fate transitions and underlying molecular changes remains a fundamental goal in developmental biology. Here, we present DeepTrack, a lineage tracing mouse model that integrates in situ cellular barcoding with high-throughput, single-cell multi-omics to simultaneously profile clonal fates, transcriptomic states, and chromatin accessibility. Using DeepTrack, we profiled clonal behaviors during gastrulation and early organogenesis, uncovered early fate priming within epiblast clones, and revealed clonal architecture within distinct regions of the nervous system. Embryo-wide multi-omic lineage tracing at single-cell resolution revealed transcriptional and epigenetic programs underlying fate commitment in neuromesodermal progenitors (NMPs). Clonal tracing with multi-omic profiles enabled inference of fate-associated gene-regulatory networks and identified the transcription factor Cdx2 as a key regulator of mesodermal specification in NMPs. Genetic perturbation of Cdx2 in chimeric embryos impaired paraxial mesoderm differentiation. Together, DeepTrack provides a versatile framework for decoding multimodal regulation of cell fate across diverse developmental contexts.

18
Accurate detection of metagenomic strain-level associations using average nucleotide identity with StrainSpy

Mallawaarachchi, S.; Tandon, K.; Rajan, N.; Marcelino, V. R.; Sandhu, S.; Bedoui, S.; Ingle, D. J.; Gunjur, A.; Tonkin-Hill, G.

2026-09-01 microbiology 10.64898/2026.08.30.748153 medRxiv
Top 2%
2.3%
Show abstract

Genetic variation among microbial strains of the same species can profoundly influence their phenotypes, ecological functions, and impacts on human health. Traditionally, the relative abundance of a species has been used to identify associations between the microbiome and disease. However, this approach overlooks intra-species genetic variation and is susceptible to spurious correlations arising from the compositional nature of abundance data and microbial load. Fast, k-mer-based algorithms can now accurately estimate strain-level Average Nucleotide Identity (ANI) in metagenomes. Despite its value as an orthogonal metric for strain-level analysis, methods for conducting ANI-based association studies remain limited. To address this, we developed StrainSpy, a statistical algorithm that identifies associations between containment ANI and variables of interest across a wide range of study designs, including longitudinal and multi-cohort designs. Re-analysis of a study examining gut microbiota recovery in 12 healthy adults following antibiotic exposure revealed novel strain-level associations, including a reduction in strain-level diversity despite species persistence. Applying StrainSpy to a multi-cohort analysis of 3,414 colorectal cancer metagenomes identified novel strain-level associations with colorectal cancer. However, in a separate collection of microbiome-immunotherapy studies, no individual strain was consistently associated across cohorts. Importantly, across both datasets, StrainSpy informed containment ANI-based machine learning models achieved comparable accuracy to traditional abundance-based methods. StrainSpy is publicly available as an R package github.com/gtonkinhill/strainspy.

19
Paired-surface spatial mechanomics links tissue stiffness maps to spatial transcriptomics

Ong, H. T.; Lou, Y.; Turley, J.; Hengst, R. M.; Ramli, M. F. H.; Shen, X.; Marlena, J.; Zhu, J.; Li, R.; Chan, C. J.; Young, J. L.

2026-08-31 bioengineering 10.64898/2026.08.29.748050 medRxiv
Top 2%
2.1%
Show abstract

Tissue mechanics influence diverse biological processes, yet directly linking stiffness measurements to spatially resolved molecular states in intact tissues remains challenging. Here we developed a paired-surface spatial mechanomics approach to map Young's modulus by nanoindentation on a fresh tissue surface and co-register the stiffness grid with 10x Genomics Visium HD spatial transcriptome bins from the immediately adjacent, parallel surface. Applied to the mouse ovary, which has spatially distinct compartments and undergoes extracellular matrix remodeling with cycle and age, the workflow generated >2,900 matched measurements across 21 regions of interest. Nanoindentation at 50-m grid spacing enabled millimeter-scale stiffness maps while balancing acquisition time in fresh tissues, with ~92 4-m transcriptome bins assigned to each stiffness value. Global and compartment-specific analyses associated stiffer regions with lower elastic fiber programs and higher inflammatory signaling, with age-dependent differences. This correlative strategy integrates experimentally measured mechanics with spatial omics in fresh tissues.

20
MetaDome 2027: a comprehensively updated resource for aggregating missense variant evidence across homologous human protein domains

Wiel, L.; Ferraro, F.; Yu, J.; Zhen, J.; Nachun, D.; Mendez, R.; Reuter, C. M.; Cui, J. L.; Bonner, D. E.; Carter, J. N.; Marwaha, S.; van de Vorst, M.; Emami, S.; Kravets, E.; Neu, M. B.; van Ham, T. W.; Kleefstra, T.; Ashley, E. A.; Bernstein, J. A.; Montgomery, S. B.; Gilissen, C.; Wheeler, M. T.

2026-08-31 bioinformatics 10.64898/2026.08.26.747388 medRxiv
Top 2%
2.1%
Show abstract

The interpretation of missense variants remains a major challenge in clinical genetics. "Meta-domains" aggregate population and pathogenic variation across homologous Pfam domain instances in the human proteome, providing per-residue context for interpreting variants of uncertain significance (VUS). Our 2019 implementation, MetaDome, is widely used and named in clinical variant-classification guidelines. Here we present the MetaDome 2027 update, featuring a comprehensively updated dataset and GRCh38 support. The redesigned pipeline enables incremental updates of GENCODE, UniProtKB/Swiss-Prot, Pfam, gnomAD, and ClinVar while maintaining 100% sequence-identity gene-to-protein mapping. Annotated Pfam domain instances grew 14.9% from 71,419 to 82,069 and meta-domain-eligible Pfam families ([≥]2 human occurrences) by 73.3% from 3,334 to 5,778; Pfam domains are annotated to 92% of human proteins. Approximately 43% of mapped protein-coding nucleotides (14.3 million in GRCh38, 13.8 million in GRCh37) are in a meta-domain; in GRCh38 67.9% (37,692 of 55,548) of pathogenic or likely pathogenic ClinVar missense variants fall at such a position. We show how MetaDome helped reclassify a de novo missense VUS in RALA and identify 52,463 ClinVar missense VUS for which meta-domains supply otherwise unavailable pathogenic evidence. MetaDome is freely available at www.metadome.app.